GOAI 2026 Global Open-source AI Challenge Grand Prize Judging Rules
To all participating teams, expert judges, and relevant organizations:
These rules are established by the organizing committee to ensure fair, impartial, and standardized judging of the Grand Prize of GOAI 2026 Global Open-source AI Challenge.
These rules govern the judging subjects, judging structure, scoring dimensions, roadshow arrangements, material requirements, scoring and calculation methods, ranking determination, result confirmation and announcement, and procedural appeals for the GOAI 2026 Grand Prize, and serve as the sole basis for Grand Prize judging and result determination.
I. General
1.1 Judging Subjects
The Grand Prize judges 6 champion teams that advanced from the four track finals:
- 1 team from Track 1 "Agent Infra"
- 1 team from Track 2 "Boundless Agents"
- 2 teams from Track 3 "AI for Research" (1 from Algorithm sub-track, 1 from Open Exploration sub-track)
- 2 teams from Track 4 "Embodied Future" (1 from Dual-arm Collaboration sub-track, 1 from All-Terrain Patrol sub-track)
1.2 Independent Scores
The finals scores of respective tracks are used only to determine track champions and are not carried into the Grand Prize. The 6 champion teams are independently re-scored and re-ranked at the Grand Prize stage.
1.3 Basic Principles
-
Objective and impartial: All advancing teams use unified evaluation dimensions, scoring standards, calculation
-
Independent judging: Judges score independently based on roadshow presentations, defense facts, and verifiable evidence. Before scores are submitted, judges shall not exchange specific scores, project rankings, or award tendencies; judge deliberation shall not replace individual independent scoring.
-
Evidence first: Scoring is based primarily on verifiable facts, including locked materials, on-site presentation, run records, code and open-source status, and experimental or hardware results. Conceptual descriptions, future plans,
-
Rule freeze: Scoring dimensions, weights, scoring formulas, tie-breaking order, recusal rules, and appeal mechanisms shall be confirmed and frozen before the event. After judging begins, substantive evaluation criteria affecting scores and final rankings shall not be temporarily added, removed, or adjusted.
-
Full traceability: Key stages including judge scoring, conflict-of-interest recusal, score verification, factual verification, dispute handling, and result confirmation shall form auditable records.
II. Judging Structure
The Grand Prize conducts cross-track horizontal judging based on the four track champions, retaining the unified finals judging framework and adding a new dimension "Scenario Deployment and Ecosystem Synergy Potential," forming six primary dimensions.
Three judging groups are established with 12 experts in total, scoring jointly by group assignment:
| Judging Group | No. of Judges | Primary Dimensions Assigned | Group Max Score | Final Weight |
|---|---|---|---|---|
| Technology Innovation Group | 4 | Innovation; Technical / Research Depth | 100 | 35% |
| Credibility and Open Ecosystem Group | 4 | Completeness and Verifiability; Open-Source Value and Reuse | 100 | 25% |
| Scenario Value | 4 | Problem Value and Real-World Impact; Scenario Deployment Ecosystem Synergy Potential | 100 | 40% |
Judges only score dimensions assigned to their group and do not cross-score other groups.
III. Scoring Dimensions and Group Assignment
| Primary Dimension | Judging Group | Evaluation Focus |
|---|---|---|
| Problem Value and Real-World Impact | Scenario Value Group | Continue to assess whether the problem is real and important, and the project's potential real-world impact |
| Innovation | Technology Innovation Group | Focus on comparing cross-track originality, breakthrough level, and relative leadership |
| Technical / Research Depth | Technology Innovation Group | Focus on technical difficulty, core moats, methodological depth, and system capability |
| Completeness and Verifiability | Credibility and Open Ecosystem Group | Assess whether core deliverables are genuinely complete and whether key claims have sufficient, credible, and verifiable evidence |
| Open-Source Value and Reuse | Credibility and Open Ecosystem Group | Assess openness, reproduction and reuse capability, and potential to continuously generate developer and ecosystem value |
| Scenario Deployment and Ecosystem Synergy Potential (New Dimension) | Scenario Value Group | Assess real-scenario entry capability and potential for synergy with industry, research, application scenarios, capital, and ecosystem resources |
IV. Scoring Sheets
All three groups use 100-point scoring sheets, each with 4 secondary scoring items. Judges enter integer scores directly.
4.1 Technology Innovation Group | 100 points (weight 35%)
| Secondary Item | Points | Key Judgment |
|---|---|---|
| Originality and Breakthrough Level | 30 | Whether substantive new methods, architectures, capabilities, or key technical breakthroughs are formed |
| Relative Leadership and Differentiation | 20 | Whether there are clear, explainable leadership points and differentiated value vs. existing papers, products, open-source solutions, or engineering practices |
| Technical Difficulty and Core Moats | 30 | Whether the core problem has sufficient technical difficulty; whether key capabilities are self-developed by the team and form identifiable technical barriers |
| Methodological Completeness and Extension Value | 20 | Whether the technical route, method, or system pipeline is rigorous and complete; whether core capabilities can extend to more tasks, scenarios, or research problems |
4.2 Credibility and Open Ecosystem Group | 100 points (weight 25%)
| Secondary Item | Points | Key Judgment |
|---|---|---|
| Deliverable Completeness | 20 | Whether core functions, research results, or system capabilities are genuinely complete and in a presentable, runnable, or verifiable state |
| Verification Adequacy and Credibility | 25 | Whether Demos, experiments, benchmarks, hardware results, code, logs, or other evidence sufficiently support key claims; whether results are verifiable |
| Openness and Reproducibility | 30 | Whether core code, models, data, tools, interfaces, or key methods have substantive open value; whether third parties can understand, run, and reproduce them |
| Reuse Value and Ecosystem Potential | 25 | Whether deliverables facilitate secondary development, research, or integration; whether they have potential for developer adoption, collaborative contribution, and ecosystem diffusion |
4.3 Scenario Value Group | 100 points (weight 40%)
| Secondary Item | Points | Key Judgment |
|---|---|---|
| Real Problem and Value Evidence | 25 | Whether a real and important problem is solved; whether user, research, industry, or social value is supported by facts and evidence |
| Actual Impact and Scaling Potential | 20 | Whether deliverables can expand usage, replicate to more scenarios, and form sustained impact |
| Real-Scenario Deployment Feasibility | 25 | What product, engineering, compliance, or delivery conditions remain to go from current Demo/deliverables to real application; whether the path is clear and feasible |
| Industry Ecosystem and Resource Synergy Potential | 30 | Alignment with industry infrastructure, research resources, application scenarios, capital, and ecosystem partners; practical potential for cooperation, POCs, incubation, or sustained development |
Note on "Industry Ecosystem and Resource Synergy Potential":
This item does not rely solely on verbal deployment commitments. It focuses on the objective alignment of the project with real resources, scenarios, and cooperation conditions, and the likelihood of substantive progress (cooperation, POCs, incubation, team or business acquisition) within the next 6-12 months.
This item does not require the project to have committed to or already deployed in a specific region; project registration or team location is not an evaluation basis.
V. Roadshow Arrangements
5.1 Time and Venue
The Grand Prize roadshow will be held on the morning of September 23, 2026, at Yunsheng Hall, Floor B1, Cloud Valley Core, with the team preparation room at Yundu Hall.
Champion teams must complete check-in and equipment checks by 7:30; the roadshow begins at 8:00. Results are
5.2 Speaking Order
Speaking order is determined by on-site draw of lots by representatives of the 6 champion teams after the finals roadshow on September 22; results are announced on-site and archived.
5.3 Format and Duration
Each team has a total presentation time of 10 minutes:
- Project presentation: 5 minutes
- Core Demo: 3 minutes
- Judge Q&A: 2 minutes
An additional 2-minute transition period is provided between teams.
5.4 Standardized Presentation Content
Each team's presentation should cover:
- Core problem addressed by the project
- Technical and product innovation
- Verification completed in the Demo
- Open-source deliverables and collaboration value
- Real-scenario or research application potential
- Next-stage development plan
5.5 Observation and Live Streaming
The Grand Prize roadshow is a controlled judging session and is not open to public observation.
Team members and necessary accompanying personnel from advancing teams may observe other teams' roadshows and public Q&A on-site throughout. Roadshow footage is streamed via official signal to the main hall, where the public may watch.
Fairness is ensured by pre-event material locking, random speaking order, uniform presentation duration, and a closed scoring process. Closing on-site observation prevents crowd impact on team performance and independent judge scoring.
VI. Material Requirements
6.1 Material Composition
The Grand Prize roadshow in principle reuses the core defense materials from track finals; champion teams are not required to create a new full PPT.
After advancing, teams may optimize narrative order on the existing materials and add 1 standardized page on "Scenario Deployment and Ecosystem Synergy" (template attached: "GOAI Supplemental Page - Deployment Path and Ecosystem Synergy").
Core materials should cover: problem and value, overall
6.2 Submission and Locking
Advancing teams must submit supplemented roadshow materials to the organizing committee by 22:00 on September 22, 2026 (see advancement notice for channel).
Materials are locked uniformly after submission. Teams may observe other roadshows and Q&A but shall not make substantive changes to locked materials based on observation or judge questions.
Teams that miss the deadline present with their track finals frozen materials.
6.3 Special Arrangements for Track 4
After entering the Grand Prize, Track 4 teams may use official finals hardware footage or Demo videos as verification materials, combined with PPT to explain core technology, system implementation, innovation, and real-scenario value. Video duration counts toward the 3-minute core Demo segment; judges may also reference September 22 finals records.
VII. Scoring Method and On-Site Execution
7.1 Each judge only completes the scoring sheet for their assigned group, entering integer scores directly, with no overall impression score.
7.2 After each team's presentation, judges complete that team's scoring immediately; after all roadshows, the score verification and result confirmation process begins.
7.3 Before scores are submitted, judges shall not exchange specific scores, project rankings, or award tendencies; they may ask questions and clarify publicly disclosed technical and factual matters.
7.4 Comments are not required. When significant abnormal score gaps, factual disputes, conflict-of-interest recusal, or committee requests for additional rationale arise, judges may add brief notes.
7.5 Scores are locked after submission. Only in cases of score entry errors, wrong material versions, or factual verification revealing errors in the original scoring basis may the judge independently adjust, with before/after records and reasons retained.
VIII. Scoring Calculation and Valid Scores
8.1 Within-Group Calculation
For the same project within the same judging group, the group score is the arithmetic mean of valid judges' totals, with no high/low scores dropped.
8.2 Valid Scores
After temporary absence or conflict-of-interest recusal, a project needs at least 3 valid scores per group to be scored; below 3, additional judging shall be arranged.
8.3 Abnormal Score Gap Review
When the gap between the highest and lowest scores within a group reaches 20 points, factual review is triggered. Review only confirms facts and evidence and does not unify opinions; after facts are confirmed, judges may maintain or independently adjust their scores.
8.4 Verifiability of Key Facts
When major disputes arise over code, Demo, open-source status, data sources, or technical attribution, materials, footage, logs, and other evidence may be reviewed. Verification addresses facts only and does not re-evaluate professional judgments.
IX. Composite Score and Ranking
9.1 Calculation Formula
Grand Prize composite score = Technology Innovation Group average × 35% + Credibility and Open Ecosystem Group average × 25% + Scenario Value Group average × 40%
All three groups use 100-point sheets; the composite score is out of 100.
9.2 Calculation Precision
Composite scores are uniformly rounded to two decimal places for display.
9.3 Ranking Determination
Grand Prize final rankings are determined by the weighted formula in Article 9.1.
Judge deliberation only confirms scoring completeness, verifies objective facts affecting scores, and handles procedural issues; it does not unify professional judgments or alter formula-derived results through deliberation, discussion, or collective voting.
No back-office role may modify judge scores. When
9.4 Result Review
After back-office score verification, a provisional result is formed. It may only be formally locked and announced after a result health check (including scoring completeness, recusal status, calculation formulas, abnormal gaps, and unresolved items) and the dispute handling window.
X. Conflicts of Interest and Recusal
10.1 Judges with any of the following relationships with a team shall declare in advance and recuse from that project:
- Employment relationship
- Investment relationship
- Supervisor / student relationship
- Collaborative R&D relationship
- Immediate family relationship
- Other material interest relationships that may affect independent judgment
10.2 No two formal judges from the same institution serve on the same judging group, to maintain diversity of evaluation sources.
10.3 Recused judges do not participate in Q&A, scoring, ranking discussions, or tie-breaking for the relevant project; missing scores are handled per Article 8.2 valid score rules.
XI. Tie-Breaking
If composite scores are tied, handle in the following order:
11.1 Firstly compare unrounded raw composite scores.
11.2 If raw composite scores remain tied, compare group averages in sequence:
- Scenario Value Group average
- Technology Innovation Group average
- Credibility and Open Ecosystem Group average
11.3 If still undetermined, non-recused Grand Prize formal judges hold an anonymous vote.
11.4 If still tied, the organizing committee organizes supplementary review before confirmation.
No new scoring indicators may be temporarily added after a tie occurs.
XII. Result Confirmation and Announcement
12.1 Two-step confirmation: Grand Prize results require two confirmations. Back-office score verification only forms a provisional result; formal locking occurs only after procedural review and the dispute handling window.
12.2 Signature and notarization: Final results are jointly signed by the judging lead, score lead, and supervision/notary personnel, and notarized on-site.
12.3 Announcement method: Grand Prize results are announced uniformly at the GOAI DAY Awards Ceremony on September 23, 2026. Before formal locking, intermediate scores and provisional rankings are not disclosed to any team, judge, or third party.
12.4 Delayed announcement: If major factual or procedural disputes cannot be reliably confirmed before the ceremony, the organizing committee may defer confirmation and announcement of the Grand Prize, while other confirmed awards and ceremony proceedings proceed normally. The organizing committee will complete
12.5 Information protection: To protect judging independence and team rights, the organizing committee does not disclose other teams' scoring records or individual judge scores.
XIII. Procedural Appeals
13.1 Acceptable Matters
Teams may file procedural appeals regarding:
- Wrong submission or material version used
- Score entry or composite score calculation errors
- Required recusal not executed
- Clearly inconsistent rule enforcement across teams
- Obvious errors in objective records or technical factual verification
13.2 Non-Acceptable Matters
Procedural appeals shall in principle not accept:
- Simple disagreement that scores are too low
- Objections to judges' professional judgments
- Requests to replace judges and rescore
- Requests for rescoring due to awards or ranking results
Differences in professional judgment are a normal part of judging. Unless factual, procedural, or conflict-of-interest
13.3 Appeal Channel and Deadline
Appeals must be filed in writing by a team representative to on-site competition supervisors within 30 minutes after Grand Prize rankings are announced, with specific reasons stated.
The organizing committee shall complete review and respond before the Awards Ceremony. If review confirms a procedural or factual error requiring correction, the organizing committee shall correct it per frozen rules and archive records.
XIV. Miscellaneous
14.1 These rules are implemented and procedurally interpreted by the GOAI organizing committee.
14.2 On-site situations not explicitly covered by these rules shall be handled per published competition rules, frozen scoring standards, the fair and consistent principle, and verifiable facts.
14.3 These rules take effect upon publication and expire upon completion of GOAI 2026 Grand Prize judging.
GOAI 2026 Global Open-source AI Challenge Organizing Committee
September 18 , 2026